The American Journal of Human Genetics
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match The American Journal of Human Genetics's content profile, based on 234 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit.
Sanchis-Juan, A.; Mostovoy, Y.; Stenton, S. L.; Ganesh, V. S.; Weisburd, B.; Yenkin, A.; Kurtas, N. E.; Zhao, X.; Shin, E.; Boone, P. M.; Su, H.; Lee, A. S.; Yadav, R.; Allan, K.; Argilli, E.; Austin-Tse, C.; Barry, B. J.; Baxter, S.; Beggs, A. H.; Bell, K. M.; Blankenmeister, B.; Bönnemann, C. G.; Brownstein, C. A.; Bujakowska, K. M.; Carbonell, E.; Cooper, S. T.; Covill, L. E.; DiTroia, S.; Donkervoort, S.; Engle, E. C.; Gallacher, L.; Genetti, C. A.; Gleeson, J. G.; Guan, B.; Hall, S.; Hildebrandt, F.; Hufnagel, R. B.; Jurgens, J. A.; Khorgade, A.; Lemire, G.; Liau, E.; Ma, J.; Madden, J.
Show abstract
Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.
Taliun, D.; Gagliano Taliun, S. A.
Show abstract
As population-scale whole-genome sequencing datasets continue to expand, they enable genetic association studies beyond single-nucleotide variants to more complex forms of genetic variation, including classical human leukocyte antigen (HLA) alleles. The HLA region comprises nine highly polymorphic classical HLA genes in extensive linkage disequilibrium that are associated with numerous autoimmune and infectious diseases. However, unlike genome-wide association studies of single-nucleotide variants, there is no general guidance for controlling the multiple-testing burden in HLA allele association analyses. Here, we systematically evaluated the effective number of independent HLA allele tests using sequencing data from diverse genetic ancestries, analytical derivation and simulations. We show that the multiple-testing burden depends on genetic ancestry, allele frequency, and the phenotype model, but remains remarkably stable across minor allele count thresholds, corresponding to approximately 60-70% of the total number of tested HLA alleles. Simulations further demonstrate that the effective number of tests can exceed 90% under realistic disease models. Analyses of 4-field HLA alleles from long-read sequencing showed that higher typing resolution increases the number of alleles but preserves the underlying correlation structure and scales the effective number of independent tests proportionally. Our results provide practical guidance for HLA association studies and support Bonferroni correction based on the total number of tested HLA alleles as a simple and robust approximation when permutation-based approaches are impractical.
Wagenknecht, J. B.; Haque, N.; Gresser, J.; Dong, X.; Zimmermann, M. T.
Show abstract
ARID1B is the most frequently de novo altered gene across a spectrum of human neurodevelopmental disorders and cancers. We found 1,456 missense variants in ARID1Bs C-terminal domains, 94% of which are clinically uninterpreted and 59% of which are observed in human disorders and cancers. Using integrative modeling of ARID1B cBAF and DNA bound structures, we calculate how these variants impact key functional sites, structural stability, and topological features, finding that 617 variants clearly impair ARID1Bs interactions within BAF, interactions with DNA, fold stability, or post-translational regulation. Our study enables scalable and precise variant interpretation by illuminating the molecular mechanisms by which missense variants damage ARID1B function, with commonalities across malignancies and congenital anomalies mutually informing therapeutic developments.
Spor, L. M.; Liau, E. M.; Sanchis-Juan, A.; Silva, A. N.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Tuitele, J.; Kershaw, E. E.; Deka, R. D.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Carlson, J. C.; Brand, H.; Minster, R. L.
Show abstract
Structural variants (SVs) are often excluded from genetic research because they are difficult to call, but they can have substantial effects on phenotypic traits. SVs have not previously been characterized in Samoans, an understudied population with a high burden of complex diseases. Using short-read whole genome sequencing data, we called SVs in 1,276 Samoans and created a Samoan-specific imputation panel inclusive of both SVs and single nucleotide variants (SNVs), called the Soifua Manuia-SV panel. Using this panel, we imputed SVs and SNVs in 3,611 Samoans with array data, enabling analysis of SV-phenotype associations in a sample of 4,887 Samoan participants. We evaluated imputation performance in Samoans against two other reference panels: (i) an SNV-only Samoan-specific reference panel, to assess whether SV inclusion impacts SNV imputation, and (ii) an SV and SNV, multi-ancestry reference panel composed of 1000 Genomes participants, which did not include Polynesians, to assess the importance of including the target population in the reference panel. The Soifua Manuia-SV panel substantially outperformed the multi-ancestry SV and SNV panel, yielding 5.5 million more high-quality (r2[≥]0.8) variants, including over 8,000 more high-quality SVs. SNV imputation based on the two Samoan-specific panels performed similarly overall, suggesting that SV inclusion does not strongly impact SNV imputation quality. This work highlights the importance of population representation for accurate imputation.
Olasege, B. S.; Campos, A. I.; Sidorenko, J.; Lin, T.; Barry, C.-J. S.; Maseras, G. T.; Vilhjalmsson, B. J.; Wray, N. R.; Hivert, V.; Yengo, L.
Show abstract
Non-random participation in genetic studies can bias associations between genetic variants and outcomes. Existing methods to detect ascertainment bias often require individual-level data, thus limiting their broad applicability. Here, we introduce a summary-statistics-based method to detect and quantify ascertainment bias in large-scale genetic studies. Our method estimates a parameter,{theta} , which captures deviations in the mean polygenic score (PGS) of an ascertained sample relative to its expectation across non-ascertained or differentially ascertained references. We show through extensive simulations that our method is robust to population stratification and reference misspecification unlike naive mean PGS comparison. When applied to 21 traits across 11 large-scale biobanks, our method recapitulates known patterns of ascertainment and detects new evidence of ascertainment on genetic susceptibility to depression, height and blood pressure in many biobanks. Overall, our framework enables systematic assessment of ascertainment directly from summary statistics and provides a scalable tool for evaluating representativeness in large scale genetic studies.
Chen, D. Z.; Mendes, M.; Ma, C.
Show abstract
The gene-sex association quality control (QC) metric was developed during the era of early small-sized genome-wide association studies and typically filter variants based on controls alone. While this practice had minimal impact in early studies, in contemporary large-scale settings they can introduce systematic bias by disproportionately discarding variants with pronounced sex differences in allele frequencies (AF), potentially removing real gene-sex interaction signals. To address this limitation, we introduce Sex-Prevalence Adjusted Allelic Difference Estimates (SPADE) metric, an X chromosome-inclusive QC framework that not only preserves potentially informative variants exhibiting gene-sex interaction but also enables a rapid and exploratory scan for such interaction. SPADE adjusts for sex-specific disease prevalence, maintains correct type I error control, and reduces the risk of falsely excluding variants compared to existing QC approaches. Extensive simulations across diverse disease architectures demonstrated that SPADE remained well calibrated, whereas the controls-only approach exhibited massively inflated type I error when variants were associated with disease. We further developed an open-source command-line software implementation to facilitate its application in large-scale genetic studies. Applying SPADE to an autism spectrum disorder (ASD) case-control cohort (6,873 cases and 8,981 controls) comprising of the Autism Speaks MSSNG, Simons Simplex Collection (SSC), and Simons Powering Autism Research (SPARK) datasets, we demonstrate that SPADE is well calibrated relative to a controls-only approach. Notably, the complementary exploratory gene-sex interaction scan implicates a sex antagonistic region encompassing RBMX2, SLC25A14, and BCORL1 (lead SNP rs150885581: A>G, p = 6.19 * 10 ^ -9), providing candidate genes for future functional investigation.
Ahn, K.; House, J. S.; Burkholder, A.; Tran, T. C.; Breeyear, J. H.; Justice, C. M.; Durney, J.; Jones, A. M.; Reyes, P. S.; Bailey, M. H.; Davis, M. F.; Vicenti, A. T.; Karnes, J. H.; Hollenbach, J. A.; Fargo, D. C.; Ginsburg, G. S.; Woychik, R. P.; Denny, J. C.; Motsinger-Reif, A. A.
Show abstract
The human leukocyte antigen (HLA) region is the strongest genetic contributor to many immune-mediated diseases, yet whether HLA architecture is shared across ancestries remains unclear. We analyzed high-resolution HLA variation in 390,823 participants from the All of Us Research Program spanning six genetic ancestry groups, including 262,915 with linked electronic health records. Using whole-genome sequencing and graph-based inference, we genotyped 20 HLA genes at G-group resolution and identified 4,780 distinct alleles. Analyses accounting for disparate sample sizes demonstrated that ancestry-private allelic variation reflected unequal discovery depth rather than ancestry-population specificity. A meta-analysis of ancestry-stratified phenome-wide association analyses with 363 HLA alleles with frequency > 0.001 and 3,430 clinical phenotypes identified 1,461 significant HLA-phenotype associations (FDR < 0.05). Although many associations reached significance in only one ancestry group, effect directions were largely concordant, highlighting differences in allele frequency, linkage disequilibrium, and statistical power among ancestry groups. Stepwise conditional modeling demonstrated that common complex trait variation could be concurrently explained by five to seven independent HLA allele signals. These findings demonstrate that a multi-ancestry, phenome-wide study can distinguish true biological heterogeneity from sampling-driven detectability differences in HLA.
Kore, P.; Tan, T.; Lu, W.; Manuel-Friedman, A.; Hu, L.; Chatterjee, N.; Zhou, W.; Dhindsa, R. S.; Atkinson, E. G.
Show abstract
Rare-variant association studies enable the discovery of high-impact genetic contributors often missed by conventional genome-wide association studies focused on common variation. However, standard burden tests aggregate variants without accounting for local ancestry in admixed genomes, reducing power when rare variant frequencies or genetic effects differ across ancestral backgrounds. Here, we introduce Tractor-Burden, an ancestry-aware gene-based association method that partitions rare-variant burden by inferred local ancestry and estimates ancestry-specific effects within a unified regression. In simulations, Tractor-Burden is well calibrated and improves power over standard burden tests under effect heterogeneity. Applied to whole-genome sequencing data from 47,152 admixed African-European individuals in the All of Us Research Program, Tractor-Burden recapitulates known associations, including ancestry-enriched effects at LDLR, and identifies additional suggestive genes and pathways for type 2 diabetes. Tractor-Burden extends rare-variant association testing to admixed genomes and provides a scalable framework for detecting and interpreting gene-level effects across local ancestry backgrounds.
Wang, X.; Wang, J.; Tiezzi, F.; Huang, Y.; Huang, W.; Maltecca, C.; Jiang, J.
Show abstract
In livestock populations, genome-wide association studies (GWAS) can produce strong, apparently localized associations even when no truly discrete nearby causal effect exists. This occurs because small effective population sizes, strong family structure, long-range linkage disequilibrium (LD), and diffuse polygenic architecture can cause the effects of many variants to accumulate and be captured jointly across broad genomic intervals, making variant-level associations difficult to interpret biologically. Using real genotypes, we constructed a benchmark in which phenotypes were simulated under diffuse polygenic architecture across a genome partitioned into alternating effect and null windows, with central-null regions (at least 1 Mb away from effect-containing regions) positioned to detect long-range LD-driven signal propagation. We evaluated nine configurations of six GWAS methods (BOLT-LMM, REGENIE, fastGWA, FarmCPU, BLINK, and SLEMM) under this architecture. The central finding is that strong associations, of the kind normally read as evidence of nearby moderate- or large-effect variants, are produced by many of these methods even though the simulated signal is distributed across many tiny effects and cannot be localized to any single variant. The methods differed sharply in the extent of locus-level spillover: several produced large numbers of genome-wide significant loci within central-null regions, whereas the full-GRM mixed-model benchmark (SLEMM) suppressed this spillover almost entirely. These results show that, under a highly polygenic architecture with livestock-like LD, GWAS tool choice has major consequences for biological interpretation. When the goal is to localize biologically meaningful signals rather than to flag association peaks that may merely reflect tiny effects accumulated through LD across a broad block, methods that control long-range LD spillover should be prioritized.
Happ, H.; Christensen, B.; Knight, S.; Novoa, A.; Isakson, D.; Nadauld, L.; Quinlan, A.; Bonkowsky, J. L.
Show abstract
Background and Objectives: Leukodystrophies are rare genetic diseases affecting the central nervous system white matter, leading to progressive disabilities and death. Although early diagnosis is critical for therapies, the penetrance and phenotypic spectrum of many leukodystrophies remain poorly defined. Here, we integrate sequencing population screening with longitudinal electronic health record (EHR) data. Our goals were to assess the prevalence of undiagnosed leukodystrophy, characterize phenotypic variability among genotype-positive individuals, and estimate penetrance across multiple leukodystrophies. Methods: We analyzed 19 genes associated with 13 leukodystrophies in pediatric and adult individuals recruited via the HerediGene Population Study, a 5-year study conducted primarily of healthy individuals in the U.S. intermountain west. Sequencing was performed on 210,983 individuals, consisting of genome sequencing for 34,033 and SNP panel imputation for 176,950. Variant results were cross-referenced to comprehensive and longitudinal (20+ years) clinical data in the Intermountain Health Enterprise Data Warehouse and to the Utah Leukodystrophy Program. Results: Pathogenic variants were identified in 4 genes (CSF1R, PLP1, POLR3A, SNORD118) in 9 individuals, none of whom had a clinical leukodystrophy diagnosis or characteristic MRI findings. These findings suggest that missed clinical diagnoses of most leukodystrophies are uncommon in a centralized healthcare system, but also demonstrate that for some leukodystrophies there may be variable or reduced penetrance, or broader phenotypic spectra than recognized. We used published incidence estimates and the observed leukodystrophy-associated genotypes to infer penetrance ranges that varied from wide for ultra-rare leukodystrophies, to tightly bounded for more prevalent conditions. Discussion: In this predominantly healthy population, we did not find any patients with leukodystrophy who had been genetically undiagnosed but then identified by sequencing. However, we identified 9 individuals with genotypes previously reported to result in leukodystrophy, but none of whom had clinical symptoms or MRI features associated with the specific leukodystrophy. Our results support a revised model in which leukodystrophies exist along a continuum of penetrance and expressivity, with implications for newborn screening, variant interpretation, and risk stratification.
Williamson, A.; Carrasco Zanini, J.; Zoodsma, M.; Koprulu, M.; Zaidi, A.; Hunt, K. A.; Taylor-Brill, E. S.; Kohleick, L.; Genes & Health Industry Consortium 1, ; Genes & Health Research Team, ; Chinnery, P. F.; Newman, W.; Finer, S.; Pietzner, M.; van Heel, D. A.; Langenberg, C.
Show abstract
Proteogenomic studies have transformed the way we derive novel insights into human biology and pathophysiology but are currently limited by their proteomic coverage and ancestral representation. Here, we integrate rare and common genetic variation with measurements of >11,000 plasma proteins on two affinity-based platforms (>6,000 targets not previously covered) in 1,535 individuals of British Bangladeshi and Pakistani ancestry of the Genes & Health cohort. We report 3,826 high-confidence common (minor allele frequency (MAF)>1.0%) protein quantitative trait loci (pQTLs), over half of which are novel and including >200 pQTLs with greater MAF in South Asians. Systematic analyses of rare (MAF<1.0%) exonic variants identify 230 gene-protein pairs and highlight the joint and distinct contributions of rare and common variants to inherited differences in protein levels. We expand analyses beyond the nuclear genome and identify 3 mitochondrial pQTLs, including a common variant in MT-RNR1, associated with lower myelin protein zero (MPZ), identifying a potential novel mechanistic link for MT-RNR1's poorly understood role in hearing loss. We create the first proteogenomic disease network in individuals of South Asian ancestry based on 384 cis-pQTL with a shared genetic disease or risk factor signal, including conditions substantially more common in South Asians, such as metabolic diseases or pregnancy-related conditions, providing insights into the underlying mechanisms. In summary, our study demonstrates the value and scientific efficiency of proteomic studies in genetically informative and understudied populations for identifying novel causes of globally relevant diseases.
Weiner, M. A.; Hiatt, L.; Ajuyah, P.; Aliyev, E.; Dashnow, H.
Show abstract
Introduction Tandem repeats (TRs), including short tandem repeats (1-6 bp motifs) and variable number tandem repeats (7+ bp motifs), have been linked to more than 50 Mendelian diseases. However, current frameworks for evaluating gene-disease relationships do not adequately address TR-specific complexities. As a result, proposed TR locus-disease relationships are often incorrectly classified, under-evaluated, or excluded entirely, limiting discovery and leading to underdiagnosis of TR disorders. Methods We developed criTRia, a scoring framework designed to accurately evaluate TR locus-disease relationships at the locus level rather than the gene level. Building on ClinGen best practices, criTRia introduces TR-specific evidence categories and reweighted scoring. We applied criTRia to curate 65 loci from STRchive, a database of disease-associated TRs. Results We compared criTRia curations with gene-level curations from nine Gene Curation Coalition (GenCC) groups. Of 65 newly scored loci, 7 had not been previously evaluated by GenCC and 17 showed significant disagreement across groups. These differences have direct implications for whether a disease is recommended for inclusion in a diagnostic gene panel. The criTRia framework also enabled curation of previously unassessed associations, bringing the total to 77 curated TR locus-disease associations and identifying four contradictory associations. Discussion By incorporating TR-specific evidence, criTRia provides a reproducible methodology for assessing TR locus-disease relationships, improving classification consistency and establishing a foundation for better integrating tandem repeats into clinical genetic medicine and providing more accurate diagnoses.
Gold, N. B.; Zouk, H.; Yeo, J.; Lipsitz, S.; Koyama, S.; Somanchi, H.; Perez, E.; Selvaraj, M. S.; O'Grady, L.; Miller, E.; Lewis, A. C. F.; Karlson, E. W.; Strong, A.; Gold, J. I.; Rehm, H. L.; Natarajan, P.; Green, R. C.
Show abstract
Importance: Genomic newborn screening (gNBS) is a potential public health intervention, but its positive predictive value (PPV) remains uncertain. Estimating the prevalence and penetrance of pathogenic and likely pathogenic (P/LP) variants in genes prioritized for screening may clarify the long-term PPV and clinical utility of gNBS. Objective: To compare ICD-based ascertainment, electronic medical record (EMR) review, and clinical assessment of genetic disorders in adults with P/LP variants in 54 genes prioritized for gNBS. Design: Two-cohort observational study with EMR review and clinical assessment in the hospital-based cohort. Setting: The U.K. Biobank (UKB) and Mass General Brigham Biobank (MGBB). Participants: 451,877 adults from the UKB and 53,371 from the MGBB, all with exome sequencing data. Exposures: P/LP variants in 54 genes prioritized through expert consensus for gNBS, in genotypes consistent with each gene's inheritance pattern. Main outcomes and measures: The primary outcome was the absolute difference in the proportion of MGBB participants identified as affected by ICD versus EMR ascertainment. Secondary outcomes included findings from clinical assessments of undiagnosed MGBB participants, corrected UKB penetrance estimates, and extrapolation to U.S.. annual birth cohorts and living adults. Results: P/LP variants were identified in 665 UKB participants (0.15%) and 82 MGBB participants (0.15%), approximately 1 in 650. In MGBB, EMR review revealed that 58/82 individuals (70.7%) were undiagnosed, although 25 of 58 (43.1%) had documented symptoms. Disease-associated ICD codes were found in 39.0% (32/82) of participants, whereas EMR review identified symptoms in 59.8% (49/82, McNemar P<.001). Applied to UKB, this correction yielded a penetrance of 28.4% (95% CI, 18.6% to 38.2%), implying that 73 to 203 participants beyond the 51 identified by ICD codes may have clinical features of disease. Extrapolated to U.S. birth cohorts, 4,900 to 5,700 newborns per year may harbor P/LP variants in these genes and survive into adulthood. Approximately 355,000 to 410,000 U.S. adults may have P/LP variants in these genes. Conclusions and relevance: Penetrance of P/LP variants in genes prioritized for gNBS is substantially higher than ICD estimates suggest. Many adults with P/LP variants are symptomatic but undiagnosed, supporting inclusion of these genes in gNBS.
Preussner, A.; Leinonen, J. T.; FinnGen, ; Pirinen, M.; Tukiainen, T.
Show abstract
Although the Y chromosome represents roughly 2% of the male genome, it is often ignored in genome-wide association studies (GWAS). Subsequently, the potential health impacts of Y-chromosomal genetic variation remain incompletely understood. To fill this gap, we performed a phenome-wide association study (PheWAS) in FinnGen across 1,426 binary and quantitative traits using Y-chromosomal variation (frequency [≥] 1%) in 104,334 genotyped men. As Y chromosome variation is prone to population stratification, we performed carefully adjusted association analyses and further examined these through kin-based validation in 19,275 female and 24,712 male 1st degree relatives. We found 121 suggestive (p < 5.6x10-3) phenotypic associations in the Y chromosome, yet none of these were strong enough to reach phenome-wide significance (p < 3.9x10-6). While only 38 associations were supported in the kin-based validation, intriguingly we found support for a previously suggested link between haplogroup I1 and coronary heart disease (CHD; OR=1.06, 95%CI=1.02-1.11, p=3.7x10-3; male validation OR=1.05; female validation OR=0.97). The I1-CHD association was detected across distinct geographical areas within Finland and was independent from Loss of Y (LOY) and the autosomal risk to CHD, proposing a link between germline Y-chromosomal variation and heart disease risk. Overall, this study presents a comprehensive phenome-wide analysis of Y-chromosomal associations, highlighting the potential relevance of Y-chromosomal variation beyond sex determination. Our findings further emphasize the need for improved capture of Y-chromosomal variants and further analyses in biobank-scale data to allow for deeper exploration of male-specific genetic architecture of complex diseases.
Fonseca, R.; Caggiano, C.; Costantino, M.; Dominguez, O.; Kenny, E.; Dahl, A.
Show abstract
Polygenic scores (PGS) are a primary output of large-scale genetic studies and are being deployed in clinical and non-clinical settings. However, current PGS assume simple additive models that ignore context-specific genetic effects, which likely reduce their accuracy and robustness. To address this, we developed PGSC, a PGS framework to incorporate locus-specific gene-context interaction effects (GxC). Simulations show PGSC is robust under the additive model and outperforms PGS in realistic settings. Using sex, age, and statin treatment status as contexts in UK Biobank, we find that PGSC outperforms PGS on average across 48 traits, with substantial improvement in some cases, such as GxSex for testosterone, GxAge for bilirubin, and GxStatins for LDL cholesterol. PGSC consistently outperforms a simple genome-wide GxC model, ampPGS, which only outperforms PGS when a context uniformly amplifies all genome-wide additive effects. Critically, PGSC improvements replicate across ancestries in the UK Biobank and in an external cohort, the Mount Sinai Million Health Discovery Program. Finally, we test robustness to log-scale phenotypes and find that ampPGS gains vanish, while the locus-specific GxC components in PGSC persist. Overall, PGSC is a simple, robust framework that demonstrates GxC effects can improve out-of-sample PGS prediction and is a step toward precision treatment.
Yang, Q.; Zou, W.-B.; Pu, N.; Li, Y.; Hu, Y.; Wang, Y.-C.; Liu, X.; Genin, E.; Masson, E.; Wang, J.; Ferec, C.; Cooper, D. N.; Li, W.; Chen, J.-M.
Show abstract
As genomic sequencing evolves beyond rare disease diagnostics toward population screening and precision medicine, clinical variant interpretation is increasingly challenged by variants whose clinical consequences depend on biological context. Current frameworks, including the ACMG/AMP guidelines, generally assign a single classification to each variant regardless of inheritance state or genetic context, potentially failing to communicate context-dependent clinical consequences. Here, we address this issue using loss-of-function variants in LPL as a uniquely informative model system in which residual physiological LPL activity can be directly quantified in vivo. By systematically integrating published biallelic LPL genotypes, physiological measurements, functional studies, and clinical phenotypes, we identified a biologically meaningful transition at approximately 10% residual physiological LPL activity. Activity below this level was predominantly associated with classical childhood-onset familial chylomicronemia syndrome (FCS), whereas higher activity was associated with phenotypic attenuation and modifier-dependent clinical expression. Furthermore, heterozygous loss-of-function variants exhibited an estimated penetrance of 5-7% for severe hypertriglyceridemia. We therefore propose a context-dependent framework in which biallelic complete- or near-complete loss-of-function genotypes are interpreted as causative for FCS, whereas heterozygous variants are interpreted as predisposing to severe hypertriglyceridemia while retaining recognition of FCS carrier status. Together, our findings demonstrate that clinical variant interpretation should integrate available biological context--including, where relevant, allelic configuration, residual biological function, and penetrance--rather than rely on the intrinsic molecular consequence of the variant alone. More broadly, this framework provides a conceptual model for interpreting variants across the continuum from Mendelian disease to genetic predisposition in the era of precision medicine.
Hwang, I.; Talbot, A.; Head, T.; Trevino, C.; Wingo, T. S.; Kotlar, A. V.
Show abstract
MotivationParent of Origin Effects (POEs), where the effect of an an allele on a phenotype differs based on maternal or paternal inheritance implicated in growth, metabolism, and neurodevelopment. Traditional tests for POEs require family data to determine parental origins of transmitted alleles. Given that such studies are expensive and time consuming compared to genome-wide association studies (GWAS), tests that function absent inheritance information are highly desirable. We develop a method, based on community detection from machine learning, that infers POEs via a spectral decomposition, obtains confidence intervals via a non-parametric bootstrap, and safeguards against confounding by non POE sources of variation. We refer to our method as Parent of Origin Inference via Spectral Estimation (POISE). ResultsWe demonstrate that POISE is well-calibrated under both Gaussian and heavy-tailed noise in simulation studies, with improved robustness to true POEs compared to existing covariance-based tests. POISE provides per-trait effect estimates with bias-corrected bootstrap confidence intervals and incorporates an information-theoretic minimum detectable effect size that filters unreliable estimates, conferring robustness to covariance-deflating variance QTL. We then apply POISE to GWAS data from the UK Biobank using BMI, LDL cholesterol, and HDL cholesterol. POISE recovers established POE loci and identifies 134 additional variants at genes implicated in lipid metabolism, immune regulation, and growth. Availability and implementationThe code for this method in Python is available at https://github.com/bystrogenomics/POISE.
Mason, A. C.; Ballabio, G.; Paz, V.; Sofat, R.; Garfield, V.
Show abstract
Mendelian randomization (MR) is widely used to infer causal relationships using genetic variants as instrumental variables, yet the selection of genetic instruments is not always given sufficient attention. Many MR studies rely on default linkage disequilibrium (LD) clumping parameters (r2 <0.001, 10,000 kb), as implemented in commonly used tools, without assessment of their suitability for specific exposures. We investigated whether this approach yields optimal instruments or whether a more pragmatic strategy yields stronger instruments. Using UK Biobank data, we examined three distinct exposure types-circulating amino acids, body mass index (BMI), and major depressive disorder (MDD). For each phenotype, we systematically varied LD clumping thresholds (r2 and genomic distance) and evaluated each instrument via both their average strength (F-statistic) and total strength (R2). Across all phenotypes, optimal instruments differed from default parameters and varied by exposure. For amino acids and BMI, more stringent LD thresholds (r2=0.00001) combined with larger clumping windows improved instrument strength, whereas for MDD, a highly polygenic, binary trait, smaller windows with stringent r2 maximized variance explained while maintaining F-statistics above the desired threshold (>10). Notably, increasing the number of SNPs did not consistently improve instrument quality, highlighting a trade-off between instrument strength and potential pleiotropy. We demonstrate that universal reliance on default LD clumping parameters can lead to suboptimal instruments. We propose a pragmatic framework for instrument selection based on empirical evaluation of strength metrics, improving the robustness and transparency of MR analyses across different exposure types.
Huang, S.; Ng, K.; Lu, Y.; Xie, Z.; Hu, C.; Zhu, B.; Zhu, W.; Lek, A.; Ma, K.; Lek, M.
Show abstract
Pathogenic variants in SGCA, encoding -sarcoglycan, cause an autosomal recessive limb-girdle muscular dystrophy, LGMDR3/2D, yet clinical interpretation of SGCA variants remains challenging due to the high prevalence of rare missense variants. -sarcoglycan is an essential component of the sarcoglycan complex at the muscle cell membrane, and pathogenic variants frequently impair its membrane localization. Here, we systematically assess the effects of all possible single-nucleotide variants across the SGCA coding sequence using a saturation mutagenesis-based experimental assay that quantifies -sarcoglycan surface expression. We generate a comprehensive functional atlas that distinguishes tolerated and damaging variants, aligning with independent genetic and clinical evidence, and reveals domain-specific properties of the cytoplasmic region, in which C-terminal truncating variants retain membrane localization, suggesting possible pathogenic mechanisms beyond impaired trafficking. This work provides a scalable functional framework to support genetic diagnosis and variant interpretation in sarcoglycanopathies. Graphical AbstractSchematic overview of saturation mutagenesis-based functional mapping of SGCA. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=187 SRC="FIGDIR/small/740214v1_ufig1.gif" ALT="Figure 1"> View larger version (41K): org.highwire.dtl.DTLVardef@1dc3d7corg.highwire.dtl.DTLVardef@48a213org.highwire.dtl.DTLVardef@88bec7org.highwire.dtl.DTLVardef@1a51bea_HPS_FORMAT_FIGEXP M_FIG C_FIG
Nielsen, M. C.; Mentzel, C. M. J.; Stoltze, U. K.; Hagen, C. M.; Baekvad-Hansen, M.; Byrjalsen, A.; Sunde, L.; Lundquist, A. A.; Lund, A. M.; Tfelt-Hansen, J.; Masmas, T.; Soerensen, E.; Pedersen, O. B. V.; Erikstrup, C.; Ostrowski, S. R.; DBDS Genomic Consortium, ; Hjalgrim, H.; Nyegaard, M.; Schmiegelow, K.; Hansen, T. v. O.; Wadt, K.; Bybjerg-Grauholm, J.; Rasmussen, S.
Show abstract
Genetic screening for rare pathogenic variants facilitates early detection and prevention of disease manifestations in medically actionable disorders, but sequencing costs limit widespread use. We introduce DoBSeq, a low-cost, high-throughput screening framework for detecting rare, single-nucleotide variants and indels. The framework includes: extraction of DNA from dried blood spots used in neonatal screening, automation of two-dimensional DNA pooling and library preparation, high-depth targeted sequencing using a 582-gene custom panel, and a probabilistic model to assign rare pathogenic variants to individuals. Benchmarked against whole-genome sequencing across 582 genes in a batch of 576 individuals, the framework detected 95% of all variants and recovered all clinically relevant pathogenic single-nucleotide variants in American College of Medical Genetics and Genomics (ACMG) actionable genes. Applied to 2304 anonymised blood donors, it yielded variant frequencies consistent with existing population estimates. At a sample cost of 29 USD, including 11 USD running costs, this framework provides a cost-efficient approach to population-level genetic screening.